Skip to content

trunk-merge/pr-108462/4f94a222-6bc6-4e2e-a33d-8e82ca5e08cd - #108634

Closed
trunk-io[bot] wants to merge 78 commits into
masterfrom
trunk-merge/pr-108462/4f94a222-6bc6-4e2e-a33d-8e82ca5e08cd
Closed

trunk-io[bot] wants to merge 78 commits into
masterfrom
trunk-merge/pr-108462/4f94a222-6bc6-4e2e-a33d-8e82ca5e08cd

Conversation

@trunk-io

@trunk-io trunk-io Bot commented Sep 29, 2026

Copy link
Copy Markdown
Trunk Merge Pull Request Banner

This pull request was created and is being managed by Trunk Merge.

This pull request is based on the master branch at SHA 9e8ba77d0cf5d60eec32aa45051949cc4c3972e1.

See more details here.

When CI completes, this pull request will be closed automatically.

Pull Requests Being Tested

This pull request is testing the changes from pull request 108462.

Dependencies

This pull request depends on the changes from pull requests 107043, 108508, 107308, 106274, 105912, 105927, 108536, and 108583.

danielcarletti and others added 30 commits September 22, 2026 17:17
…ngress

The Postgres CDC adapter writes cdc_ingest_mode=buffered with the slot and
publication fields, so a source that turns on CDC never runs the legacy lane.
Every streaming schema whose first sync is done now serves the buffer in any
table mode; the per-schema cdc_buffered_lane marker is no longer read, so a
table added after a flip joins without a re-run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Repair CDC and the automatic lost-slot recovery both recreate the slot through
the Postgres adapter's recreate_slot, after resetting every schema to snapshot.
Nothing from the dead slot is owed to the legacy lane, so the new slot starts
buffered, as a new source does. The snapshot-to-streaming purge already clears
any stale buffer files before the consumer reads them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
On a buffered source, a table still taking its snapshot now has its changes
written to the S3 buffer instead of legacy deferred runs. The snapshot run
stamps when it started reading, and the hand-over to streaming deletes only the
buffer files from before that stamp. The consumer then replays every later
change over the snapshot, which converges because the merge is an upsert by
primary key.

A TRUNCATE reset now drops the run's pending changes for the table and purges
its buffer strictly, so no change from before the TRUNCATE survives to bring
back rows. Legacy sources keep the full purge at hand-over.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…efore advancing

A TRUNCATE is now handled before every slot advance, so a failed reset or
purge fails the run while the slot still holds the TRUNCATE and the retry
repeats it. Capturing a snapshotting table to the buffer is gated by the
dwh-cdc-buffered-snapshot flag, to be turned on once no worker runs the
previous release, which would otherwise defer the same table's changes or purge
its whole buffer at hand-over. The batcher records each table's byte estimate
when it adds an event, so a discard releases exactly what the table held.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Repair resets only the CDC tables with sync on, but the source now comes back
buffered, so a table with sync off could later consume stale files from before
the repair without a new snapshot. Repair now purges the buffer of every CDC
table on the source. The lane predicate docstrings and the setup comment are
corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the flip command

New sources skip the flip command, so its preflights have to hold elsewhere.
The loader now always resolves write ordering for CDC batches instead of
reading the dwh-cdc-write-resolution flag, which could fail closed. A table
whose source has a _ph_cdc_seq column is refused when a source is created or
a table is switched to CDC, rather than failing the whole source on its first
sync. The unused cdc_buffered_lane marker and the duplicate lane predicate are
removed, and the flip command no longer writes the marker or re-runs on an
already-buffered source. Repair also resets CDC tables with sync off, so a
table turned back on re-snapshots instead of streaming past the lost gap.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… into claude/cdc-buffered-snapshot-handoff

# Conflicts:
#	products/warehouse_sources/backend/temporal/data_imports/cdc/source_manager.py
…-snapshot-handoff

# Conflicts:
#	products/warehouse_sources/backend/management/commands/migrate_cdc_source_to_buffered.py
#	products/warehouse_sources/backend/temporal/data_imports/cdc/source_manager.py
#	products/warehouse_sources/backend/temporal/data_imports/cdc/tests/test_source_manager.py
…-snapshot-handoff

Restores repair.py to master's version: the previous master merge kept
pre-squash hunks of the base PR (a second strict purge after the new
slot exists).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…keep its buffer

A snapshot whose changes the buffer carries now keeps an unbroken run of
them, and the hand-over replays all of it instead of deleting files older
than a time cutoff. A schema marker (cdc_snapshot_lane) records the
decision, so routing no longer follows the flag on every run: one
snapshot cannot split between the buffer and deferred runs when the flag
flips or fails. Without the marker the hand-over purges the whole buffer,
as legacy did.

- capture starts a snapshot in the buffer only with no deferred runs,
  and empties its buffer first so files from before a gap never replay
- resets of a table the buffer already serves keep it there
- the decoder exposes a truncate only after its transaction's changes
  are consumed, so a mid-run purge never runs ahead of them
- a truncate purge is strict on any buffered source
- the rollback command refuses while a snapshot runs in the buffer
- drops the snapshot start stamp and the time-bounded purge

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…napshot

A TRUNCATE or a lost slot resets the table to a snapshot. The reset set
the buffer marker before the strict purge, so a failed purge left the
marker over pre-reset files. The hand-over keeps every file of a marked
table, so a snapshot that finished before capture retried would replay
those files and bring truncated rows back. Purge first, so a failed
purge leaves the table unmarked and the next run repeats the reset.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…hand-over

The hand-over read the snapshot marker, purged the buffer, then flipped
the table to streaming under a separate lock. A capture run could mark
the table and write its first file between the read and the purge, and
the strict purge then deleted changes the slot had already released.
The hand-over now holds the row lock from the marker read through the
flip, and capture sets the marker only while the row still snapshots,
so a mark after the flip cannot outlive the snapshot.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Capture now only writes the S3 change buffer, and each table's own
scheduled sync loads it. Leftover legacy state converts itself before
capture reads the WAL: a legacy source gets its buffer emptied and its
table schedules rebuilt, a table with deferred runs snapshots again in
the buffer, and capture job rows left running are failed.

Removes capture's own delivery (write trackers, deferred runs, the
backpressure guard, the shadow lane), CDCHandledExternally, the flip and
shadow-validation commands, and the two rollout flag checks. The
stalled-schedule sweep now reports streaming CDC tables whose consumer
stopped.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ffer

The schema resync endpoint reset a streaming table to snapshot without
the snapshot marker, unlike the table-mode path. The next capture run
then treated it as a new snapshot and emptied the table's buffer,
deleting any changes an in-flight capture run had written after the
snapshot started reading. Resync now sets the marker whenever the table
is already streaming from the buffer, as the table-mode path does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…shots

- Confirm the slot past the decoder's last commit at the end of a run,
  so a TRUNCATE in a transaction of its own is never read and reset
  twice. A reset also cancels the table's running sync, so a snapshot
  that began before a repeated reset cannot hand over without the
  changes the reset drops.
- Drop the snapshot marker when sync is turned off, and when a table
  joins the capture set again: the buffer has a gap by then.
- Admin trigger-sync resets and resync_schemas_non_billable set the
  marker like the resync endpoint, and write through the locked merge.
- Stop removing the marker from a request's stale copy in the resync
  and table-mode resets.
- update_sync_type_config_for_reset_pipeline and
  set_partitioning_enabled merge under the row lock instead of saving
  the copy the sync loaded, which could wipe a marker capture set.
- Rollback restarts snapshots the buffer carries on legacy instead of
  refusing until they finish.
- Capture's own unmarked reset stops buffering the table for the rest
  of the run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-snapshot-handoff

# Conflicts:
#	products/warehouse_sources/backend/temporal/data_imports/cdc/activities.py
…ollback restart

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…emove-legacy-lane

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… restart

The conversion that restarts a deferred-run table's snapshot now pauses the
table's schedule, raises on a failed cancel, and resets only once the old
sync's workflow has closed and none of its batches are still loading. Until
then capture leaves the table out. A failed schedule rebuild after the reset
stays pending and every capture run retries it.

The resnapshot predicate checks the source's ingest mode and pending deferred
runs again, so a table whose buffer can hold delivered copies or a gap gets no
marker. A CDC table's sync frequency is capped at 7 days, because the buffer
expires captured changes after 14 days.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…egacy-lane

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A query can now choose the native JSON events table or the legacy one
without flipping the instance settings. When the modifier is unset, the
settings decide as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…pture converts it

Its buffer holds copies of changes the legacy lane already delivered, so a
scheduled run that a sync frequency change unpaused no-ops the tick until
conversion empties the buffer and marks the source buffered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ticking

Capture marked every table Failed on a retryable attempt, and nothing cleared
that once a later attempt succeeded, because the table's status belongs to its
scheduled sync. It now marks tables Failed only when the failure is terminal,
matching the failure job rows and digest.

A restarted legacy snapshot whose schedule rebuild was skipped by a
deliberate hold dropped its pending marker, so no later run restarted it.
The marker now stays until the rebuild happens.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s terminal

activity.info() raises outside a Temporal activity, so the failure handler
raised before it recorded the failure. Nothing retries there, so the failure
is terminal and the schema is marked failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
danielcarletti and others added 27 commits September 29, 2026 08:54
…on their old cadence

A converted legacy table had no run that ever listed the buffer, so the buffer
expiry check took it for expired and its first sync re-snapshotted it. The
conversion now stamps each table, and the check counts a recent stamp as a read.

The conversion also speeds a table set slower than its source's fastest table up
to it, the cadence the legacy lane delivered at, and it hands a deferred-run
table's snapshot restart to capture's pending-reset path instead of running its
own. A deferred-run table now joins the buffered set, so a reset that completes
before the read no longer drops that run's changes behind a gap-free marker.

The schedule picker stops offering CDC tables Monthly, which the API rejects.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The MCP server cached a token's scopes for the whole 7-day session cache
lifetime and never re-read them. A user who added a missing scope to a
personal API key stayed blocked, and reconnecting the client did not help
because the cache is keyed by the token.

Refresh the entry after 2 minutes, and keep the cached scopes when the
refresh call fails so a transient API failure does not end a live session.
The missing-scope messages now name the recovery step for a personal API
key instead of only the OAuth reconnect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: caab3a4d-6d54-4778-bc47-537abe15c716
The rejected personal key read never reaches OAuth introspection, so the
mock for it was dead. Move the session-survival comment onto the line it
explains.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: caab3a4d-6d54-4778-bc47-537abe15c716
OAuth tokens get a new value when their scopes change, so the token-keyed
cache already misses for them. The missing-scope error and search hint now
give only the recovery step for the connection's credential type.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Django's migration writer serializes a plain Enum member default as an import
of the enum's module, so the next migration that touches this field would
import the facade enums module path. Migration 0013 already records "off", so
the value default matches it and makemigrations detects no change.
enrichment_label_batch walked every org's latest fetch in
organization_id order, so a daily --limit spent itself on the oldest
archive rows and new signups waited behind them. Walk newest first by
(fetched_at, id), add --lookback-days (the Dagster job passes 14), and
share the candidate query with count_pending_candidates. That count
excluded labeled fetches before picking each org's latest one, so it
over-counted a re-enriched org whose latest fetch already had a label.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…g/fixreplay-vision-stop-search-results-00dc64

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Fold the lookback into recent_latest_fetches_qs, so the command and
count_pending_candidates share one helper. State that the window also
bounds score repairs. Add a tie case for the keyset, pin the count's
window in the job test, and give two test_ai_pilled tests the fetched_at
order their scenarios need: under the new order they passed even with
their bugs put back.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…/cdc-remove-legacy-lane

# Conflicts:
#	products/warehouse_sources/backend/management/commands/migrate_cdc_source_to_buffered.py
#	products/warehouse_sources/backend/management/commands/validate_cdc_buffer.py
#	products/warehouse_sources/backend/temporal/data_imports/cdc/activities.py
#	products/warehouse_sources/backend/temporal/data_imports/pipelines/README.md
#	products/warehouse_sources/backend/tests/management/test_validate_cdc_buffer.py
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Make both join branches of implementation_pr_report_filter team-first subqueries. The PR artefact branch matches pull_request_id = ANY(ARRAY(team PR ids)), and the assignment branch becomes an id__in subquery. Postgres then probes the artefact index with only the team's PR ids, and evaluates each branch once.

Refs #108503

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ac91e29c-71c8-47bc-9c32-7e5bfe19cf0b
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: ac91e29c-71c8-47bc-9c32-7e5bfe19cf0b
@trunk-io trunk-io Bot closed this Sep 29, 2026
@trunk-io
trunk-io Bot deleted the trunk-merge/pr-108462/4f94a222-6bc6-4e2e-a33d-8e82ca5e08cd branch September 29, 2026 18:42
@trunk-io

trunk-io Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Author

Static Badge   Static Badge   Static Badge

Failed Test Failure Summary Logs
Lemon UI/Object Tags EditableTagsInNarrowContainer play-test The test failed because the expected narrow tags container was not found. Logs ↗︎

View Full Report ↗︎ ⋅ Docs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants